Speaker recognition by separating phonetic space and speaker space
نویسندگان
چکیده
In speaker recognition, it is a problem that speech f e a-ture varies depending on sentences and time diierence. This variation is mainly attributed to the variation of phonetic information and speaker information included in speech data. If these two kinds of information are separated each other, robust speaker recognition will be realized. In this study, w e propose a speaker identiica-tion and speaker veriication method by separating the phonetic information and speaker information by a sub-space method, under the assumption that a space with large within-speaker variance is a \phonetic space" and a space with small within-speaker variance is a \speaker space". We carried out comparative experiments of the proposed method with a conventional method based on GMM in an observation space as well as in a space transformed by L D A. As a result, we could construct a robust speaker model with a few model parameters using a few training data by the proposed method.
منابع مشابه
Eurospeech 2001 - Scandinavia SPEAKER RECOGNITION BY SEPARATING PHONETIC SPACE AND SPEAKER SPACE
In speaker recognition, it is a problem that speech f e a-ture varies depending on sentences and time diierence. This variation is mainly attributed to the variation of phonetic information and speaker information included in speech data. If these two kinds of information are separated each other, robust speaker recognition will be realized. In this study, w e propose a speaker identiica-tion a...
متن کاملGenerative Acoustic-Phonemic-Speaker Model Based on Three-Way Restricted Boltzmann Machine
In this paper, we argue the way of modeling speech signals based on three-way restricted Boltzmann machine (3WRBM) for separating phonetic-related information and speaker-related information from an observed signal automatically. The proposed model is an energy-based probabilistic model that includes three-way potentials of three variables: acoustic features, latent phonetic features, and speak...
متن کاملSpeaker adaptation by modeling the speaker variation in a continuous speech recognition system
A method for unsupervised instantaneous speaker adaptation is presented and evaluated on a continuous speech recognition task in a man-machine dialogue system. The method is based on modeling of the systematic speaker variation. The variation is modeled by a low-dimensional speaker space and the classification of speech segments is conditioned by the position in the speaker space. Because the e...
متن کاملNoise and Metadata Sensitive Bottleneck Features for Improving Speaker Recognition with Non-Native Speech Input
Recently, text independent speaker recognition systems with phonetically-aware DNNs, which allow the comparison among different speakers with “soft-aligned” phonetic content, have significantly outperformed standard i-vector based systems [912]. However, when applied to speaker recognition on a nonnative spontaneous corpus, DNN-based speaker recognition does not show its superior performance du...
متن کاملPhysiologically-motivated Feature Extraction Methods for Speaker Recognition
PHYSIOLOGICALLY-MOTIVATED FEATURE EXTRACTION METHODS FOR SPEAKER RECOGNITION Jianglin Wang, B.S., M.S. Marquette University, 2013 Speaker recognition has received a great deal of attention from the speech community, and significant gains in robustness and accuracy have been obtained over the past decade. However, the features used for identification are still primarily representations of overal...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2001